Back

Frontiers in Systems Biology

Frontiers Media SA

Preprints posted in the last 90 days, ranked by how well they match Frontiers in Systems Biology's content profile, based on 10 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Large-scale, interpretable gene regulatory network inference through biologically informed matrix factorization

Micheletti, S.; Fanfani, V.; Vogt, J.; Quackenbush, J.; Fischer, J.; Marx, A.; Mandros, P.

2026-07-16 systems biology 10.64898/2026.07.15.738791 medRxiv
Top 0.1%
3.1%
Show abstract

Gene regulatory networks (GRNs) provide a mechanistic framework for understanding how transcription factors coordinate gene expression to establish cellular identity and phenotype. Methods that integrate gene expression with motif-derived regulatory priors and other sources of biological information have substantially advanced gene regulatory network inference by reconstructing condition-specific regulatory architecture. These approaches estimate the evidence supporting regulatory interactions and have proven remarkably successful in a wide range of biological applications. A complementary view of regulatory networks, however, seeks to estimate the effect of those interactions on gene expression itself, providing a framework in which regulatory edges can be interpreted as activating or inhibitory influences on transcription. We developed Giraffe, a biologically informed matrix factorization framework that jointly estimates transcription factor activities and gene regulatory networks by integrating gene expression, motif-based regulatory priors, and transcription factor protein-protein interactions. Giraffe estimates signed partial regulatory effects whose magnitude and sign can be interpreted as the strength and direction of transcriptional regulation. Building directly on the biological framework established by methods such as PANDA, Giraffe provides a complementary representation of gene regulatory networks that emphasizes mechanistic interpretation while remaining scalable, flexible, and computationally efficient. Across synthetic benchmarks, six human tissues, yeast transcription factor perturbation experiments, and liver hepatocellular carcinoma, Giraffe accurately reconstructs regulatory interactions while distinguishing activating from inhibitory regulation with high accuracy. The inferred networks recover known features of tissue-specific regulation, correctly classify regulatory effects in transcription factor perturbation experiments, and identify biologically coherent changes in regulatory programs associated with liver cancer. Together, these results demonstrate that estimating the direction of transcriptional regulation provides a complementary perspective on gene regulatory networks that facilitates biological interpretation and hypothesis generation.

2
Textural features for pathway-level representation of omics data in biological networks

Alexeyenko, A.

2026-07-17 systems biology 10.64898/2026.07.12.737672 medRxiv
Top 0.1%
2.6%
Show abstract

More than 50 years ago, Haralick and co-authors proposed a family of gray-level co-occurrence statistics that became known as textural features. These features are widely used in image analysis, but their application to biological networks has remained limited because cellular networks are sparse, irregular graphs rather than regular pixel grids. This work presents a network-adapted version of Haralick texture analysis for generating pathway-level features from gene-level omics profiles. The resulting profiles reduce dimensionality and can be used as candidate predictors of anti-cancer drug response. Performance of these features is compared with original gene expression variables and with pathway features from network enrichment analysis (NEA), whose robustness has been demonstrated previously. Although technically simpler than NEA, Haralick features showed comparable sensitivity. More importantly, selected Haralick features were preserved between in vitro drug screens and clinical treatment-associated survival analyses, supporting their potential use for prioritizing robust pathway-level drug-response correlates.

3
Overinflation and overconcentration: why Cauchy perturbation kernels are the right choice for ABC-SMC

Sturrock, M.; Shahrezaei, V.

2026-07-09 systems biology 10.64898/2026.06.24.734205 medRxiv
Top 0.1%
2.1%
Show abstract

Approximate Bayesian computation sequential Monte Carlo (ABC-SMC) propagates its particles with a perturbation kernel, and with the standard Normal kernel it degrades sharply as the parameter dimension grows, a failure usually attributed to dimension itself. We show instead that it is governed by the quality of the summary statistics, with dimension entering only through a separate and milder mechanism, and that the two must act together for the Normal kernel to break. The first ingredient is covariance overinflation: the kernel covariance, estimated from the particle cloud, overshoots the true posterior covariance by a factor set by information loss in the summary statistics. We derive this overscaling factor in closed form for a Gaussian model with sufficient statistics and show that it stays modest at any dimension, shrinking toward its baseline value as the tolerance tightens; the extreme values seen in practice (of order 103) are a signature of insufficient summaries, not of dimension. The second ingredient is perturbation overconcentration: the normalised Normal step size concentrates around one as the dimension grows, so every proposal overshoots by the same factor. Either ingredient alone is harmless; only their combination breaks the Normal kernel. A Cauchy kernel (multivariate t with one degree of freedom) removes the concentration, keeping a positive acceptance rate under arbitrary overscaling at a bounded worst-case cost of 1.87x in expected squared jump distance. In a Metropolis-Hastings framework we derive closed-form acceptance rates for both kernels that illustrate the advantage of the Cauchy kernel in this limit. A series of full ABC-SMC computational experiments on five problems at d = 12, including a hierarchical gene-expression model, show the Cauchy reducing the sliced Wasserstein distance to the reference posterior by factors of up to 50 with the same simulation budget. Since the summary statistics are commonly insufficient for the models that require ABC, overinflation is structural and the Cauchy perturbation kernel is the right default for problems in higher dimensions.

4
Weak form Scientific Machine Learning for Systems Biology: A Tutorial on WENDy

Heitzman-Breen, N.; Lyons, R.; Jain, P.; Jolly, M. K.; Bortz, D. M.

2026-07-09 systems biology 10.64898/2026.07.02.735880 medRxiv
Top 0.1%
1.7%
Show abstract

Mechanistic ordinary differential equation models are widely used in systems biology to represent biochemical networks, population dynamics, cell-state transitions, and other biological processes; however, their predictive value depends critically on accurate parameter estimation from noisy and often sparse experimental data. In this tutorial, we present the Weak-form Estimation of Nonlinear Dynamics (WENDy) method as a forward-solver-free approach that reformulates parameter estimation as a covariance-corrected weak-form regression problem by integrating the model equations against compactly supported test functions. We present the background on the methodology through the lens of the familiar logistic equation, and we demonstrate applications of the method on real experimental data through two systems biology examples: a glycolytic oscillator with relatively dense time-course data and a sparse epithelial-mesenchymal cellstate transition model with multiple experimental replicates. Ultimately, using WENDy, we estimate interpretable biological parameters with uncertainty for systems with noisy and sometimes sparse available experimental data.

5
Circadian Oscillation Detection Analysis and Comparison (CODAC): a Multicriteria Method to Estimate and Compare Rhythmicity

da Silveira, T. P.; Lincoln, K.; Nguyen, T.; de Assis, L. V. M.

2026-08-21 systems biology 10.64898/2026.08.17.745071 medRxiv
Top 0.1%
1.4%
Show abstract

Analysis of circadian patterns in time-series data requires computational methods that can accommodate several factors, including variable sampling resolution, replicate number, and missing values. Most existing tools simplify rhythmicity to a strict dichotomy based solely on a single p-value threshold. This leads to a level of uncertainty that affects many biological targets. We developed CODAC (Circadian Oscillation Detection Analysis and Comparison), a framework that integrates nonlinear constrained optimization with a multicriteria rhythmicity classification scheme to evaluate rhythmic patterns without relying on a single statistical cutoff. This approach allows CODAC to identify and exclude medium-confidence rhythms rather than force them into a rhythmic/arrhythmic dichotomy. CODAC comprises four modules: (i) CODAC_single estimates rhythmicity within a single group; (ii) CODAC_flex extends this to identify distinct waveform types within one group; (iii) CODAC_compare performs pairwise comparisons across two or more groups to detect rhythmic or arrhythmic changes; and (iv) CODAC_multi handles more complex designs involving multiple-group comparisons. Using in silico simulations and public transcriptomic datasets, we show that CODAC performs comparably to established methods while providing additional flexibility for rhythm classification and comparison. Taken together, CODAC provides a flexible and open-source package for circadian timeseries analysis with automated visualization tools.

6
Intricate Dynamical Cross-Talk Between p53 Protein and Cell Cycle Regulators Governs Mammalian Cell Fate

Charan, K.; Kar, S.

2026-06-10 systems biology 10.64898/2026.06.07.730771 medRxiv
Top 0.1%
1.3%
Show abstract

In mammalian cells, under normal circumstances, the p53 protein exhibits oscillatory dynamics in response to DNA damage and maintains the cells in a cell-cycle-arrested state. Intriguingly, some cells can escape this cell-cycle-arrested state even after prolonged DNA damage, and often undergo mitotic catastrophe. In this context, the precise role of p53 dynamics and its complex interplay with cell-cycle regulation remain poorly understood. Herein, by constructing a comprehensive network model, we have identified crucial crosstalk regulations between the p53 protein and key cell-cycle regulators that enable some cells to escape cell-cycle arrest during prolonged DNA damage. The model further illustrates a probable cellular mechanism underlying mitotic catastrophe and predicts ways to induce it in a therapeutically relevant manner.

7
Port of Protein-Protein Interactomes: An experiment-based protein-protein interactome database for rice

Liu, X.; Lu, J.; Jia, L.; Xia, D.; Huang, J.; Cheng, Y.; Li, M.; Chen, Y.; Liu, X.; Li, G.; Liu, W.; Li, J.; Ying, J.; Wang, Y.; Li, Z.; Tong, X.; Hou, Y.; Zhiguo, E.; Zhang, J.; Zhang, J.

2026-08-20 systems biology 10.64898/2026.08.16.744343 medRxiv
Top 0.1%
1.3%
Show abstract

Protein-protein interactions (PPIs) play a crucial role in enabling proteins to carry out their functions within various biological processes (Hui et al., 2003). Since the introduction of the yeast two-hybrid (Y2H) method for PPI detection in 1989 (Fields and Song, 1989), the identification of PPIs has become a significant focus in modern biological research. PPI goes beyond examining individual proteins, allowing researchers to establish a comprehensive network that regulates biological processes. Rice, as a key model organism in plant biological studies, has been at the forefront of PPI research. In 2008, prominent rice scientists in China called for concerted efforts to define a comprehensive protein-protein interaction network experimentally, which aimed to facilitate the prediction of the functional mechanisms operating throughout a plants lifecycle (Zhang et al., 2008). With efforts for 2 decades, the experimentally identified rice PPIs have reached over ten thousand. Several public databases have been established to systematically collate and store PPIs, including STRING (Szklarczyk et al., 2019), BioGRID (Oughtred et al., 2020), IntAct (del Toro et al., 2022), PRIN (Gu et al., 2011), RicePPINet (Liu et al., 2017) and RiceNet v2 (Lee et al., 2015). However, most PPI datasets in rice stem from computational predictions, while experiment-based rice PPI datasets are fragmented due to the lack of systematic profiling at the rice PPIome level, which largely hinders information sharing in the rice research community. To bridge this gap, we constructed the Port of Protein-Protein Interactomes (POPPIN; https://riceome.hzau.edu.cn/poppin/), an integrated database dedicated to sharing experimentally verified PPIs and functional clues in rice. Empowered by high-throughput PPIome profiling technologies and text mining assisted by a large language model (Huang et al., 2025; Liu et al., 2025), POPPIN currently has deposited over 150,451 pieces of rice PPI-related information. Additionally, POPPIN provides detailed protein information, including GO annotations, subcellular localizations, domains, trait ontology (TO) information, and hyperlinks to external biological databases. Through offering a user-friendly web interface for search and dynamic network visualization, POPPIN serves as the first large-scale, experiment-based database for searchable PPIs in rice, and has the potential to be extended to other species under this structural framework.

8
Large-scale analysis of optimisation methods for parameter estimation problems in the life sciences

Grein, S.; Penas, D. R.; Weindl, D.; Lakrisenko, P.; Banga, J. R.; Hasenauer, J.

2026-07-13 systems biology 10.64898/2026.07.11.737731 medRxiv
Top 0.1%
1.1%
Show abstract

Dynamic models are central to the computational life sciences but typically contain unknown parameters that must be inferred from experimental data. High-throughput measurements have made this task increasingly challenging, yielding high-dimensional search spaces and non-convex objectives with many local optima. This makes the choice of optimisation method critical. However, existing empirical studies either consider only a limited number of benchmark problems or only a narrow spectrum of local, global and hybrid optimisation methods. Here, we present a comprehensive benchmark of a broad range of optimisation methods on a curated collection of parameter estimation problems, comprising 990 method-problem-pairs executed on two independent supercomputing infrastructures. Our evaluation quantifies success rates, solution quality and computational cost, revealing characteristic strengths and limitations of each approach. We find that optimisation methods separated into clear performance tiers. Building on these results, we implemented a new hybrid strategy that combines enhanced scatter search with the best-performing local solver, which showed robust performance and improved on the other scatter-search variants we tested. Our results provide practical guidance for selecting optimisation methods and thereby support more accurate and reliable model calibration.

9
GreenSloth: a curated database and executable platform for mechanistic photosynthesis models

Corvest, E.; van Aalst, M.; Nies, T.; Nguyen, Q. H.; Ebeling, J.; Strauch, M.; Cisse, E.-H. M.; Hassan, T.; Matuszynska, A.

2026-07-26 systems biology 10.64898/2026.07.22.740007 medRxiv
Top 0.1%
1.1%
Show abstract

Mechanistic models of photosynthesis have expanded substantially over the past decades, covering processes from light reactions to carbon fixation. However, these models remain fragmented across the literature, inconsistently implemented, and difficult to reproduce or reuse, limiting their adoption beyond the research group that developed them. Here, we present GreenSloth, a freely accessible web-based database of 22 published mechanistic photosynthesis models, reimplemented as standardized, executable Python objects within MxlPy, an open-source framework for mechanistic biological modeling. Although the database is primarily designed for dynamic mechanistic models formulated as ordinary differential equations, the current implementation also includes the fields most widely cited steady-state mechanistic model and its variants. GreenSloth provides a structured environment for model discovery, comparison, and reuse, addressing reproducibility challenges in the field and enabling integration into emerging hybrid modeling approaches. It is also interactive: each model runs directly in the browser, with no installation, environment setup, or programming required. The resource is openly accessible and designed for long-term community maintenance, hoping to position itself as foundational infrastructure for the photosynthesis modeling community. Database URLhttps://greensloth.rwth-aachen.de/

10
Dynamics of a Hes1-Dll1 regulatory network in coupled muscle stem cells: stability, bifurcations, and coexistence of oscillatory states

Bujtar, Z.; Goldenbogen, B.; Wolf, J.

2026-07-30 systems biology 10.64898/2026.07.29.741250 medRxiv
Top 0.1%
1.1%
Show abstract

Muscle regeneration relies on the coordinated activation of muscle stem cells, whose fate decisions are regulated by intracellular gene expression dynamics and intercellular coupling via the Notch-Dll1 signaling pathway. Central components of this regulatory network include the transcriptional repressor Hes1, its target gene Dll1, and the myogenic regulator MyoD. Experimental and theoretical studies have shown that proliferating muscle stem cells exhibit oscillatory dynamics of these molecules, whereas sustained expression is associated with differentiation. Here, we investigate the dynamics of a previously established delay differential equation model of two coupled muscle stem cells. Using linear stability analysis, we systematically characterize how model parameters affect the transition between stable and unstable steady states. In addition, numerical bifurcation analysis is employed to study the influence of intercellular coupling strength and delay on the system dynamics. Our analysis shows that continuous variation of the coupling delay induces repetitive changes in the stability of the steady state. However, this sensitivity towards the coupling delay is confined to a narrow region of parameter space and therefore requires a fine tuning of all other parameters. Beside the identification of parameter sets for in-phase and out-of-phase oscillations, we demonstrate the possibility of coexisting stable in-phase and out-of-phase oscillations, a dynamical feature that has not been reported previously. While oscillation periods are largely determined by intracellular regulatory mechanisms, oscillation amplitudes can be strongly modulated by intercellular coupling. These results provide new insight into how the intracellular network and intercellular communication interact to generate different collective dynamics.

11
HetNetEX: Exact Asymptotic Inference in Heterogeneous Biomedical Knowledge Graphs

Ghosh, T.; Gillenwater, L. A.; Greene, C. S.; Costello, J. C.

2026-07-10 systems biology 10.64898/2026.07.05.736581 medRxiv
Top 0.1%
1.1%
Show abstract

Heterogeneous biomedical knowledge networks (hetnets) integrate disparate data types, drugs, genes, diseases, and pathways, across independent sources; Hetionet (https://het.io) is a widely used example. A standard approach for assessing connectivity significance is XSwap, which permutes the hetnet P times and fits a gamma-hurdle null model to the degree-weighted path count (DWPC), pooling permuted values across pairs with matching source and target degrees to increase the effective sample size. This permutation approach has been highly successful in practice, but it faces four practical constraints in large graphs: (1) a finite resolution for the smallest reportable p-values, (2) computational cost that grows prohibitive at path lengths L [≥] 4 or 5, (3) a variance model (Var {propto} {micro}2) that departs from the configuration-model form (1 +{kappa} ){micro}, and (4) O(P 10m L) runtime. To complement this approach, we present HetNetEX (Heterogeneous Network EXact inference), which computes the null DWPC distribution analytically from degree sequences using the configuration model in O(Ln) time. In simulations at P = 200 across L = 1-4, HetNetEX achieves Spearman{rho} > 0.96 concordance with XSwap rankings while being >10,000x faster and providing analytical p-values without a resolution ceiling. High-degree pairs show larger XSwap sampling error than low-degree pairs, reflecting the finite-sample nature of permutation that analytical computation avoids.

12
VFB-MCP: Natural-Language Access to Drosophila Neuroscience Grounded by an Expert-Curated Ontology-Led Knowledgebase

McLachlan, A. D.; Court, R.; Pilgrim, C.; Longden, K.; Brown, N. H. D.; Osumi-Sutherland, D.; Jefferis, G. S. X. E.; Armstrong, D. J.

2026-06-21 neuroscience 10.64898/2026.06.16.732577 medRxiv
Top 0.1%
1.1%
Show abstract

Biological databases store curated knowledge that researchers traditionally access through web interfaces or APIs. To move beyond casual browsing requires domain-specific knowledge and expertise to frame the queries necessary to explore this data. This generates a barrier for new users in scientific fields undergoing paradigm shifts. Exposing these databases to large language models (LLMs) via the Model Context Protocol (MCP) enables natural-language access, a potential accessibility solution. We implement this for Virtual Fly Brain (VFB), an expert-curated and ontology-backed knowledgebase of Drosophila neuroscience, providing the precision needed to make recently-integrated connectomes accessible. Benchmarked on 30 neuroscience tasks against a bare LLM and a web-search-assisted LLM, the VFB-MCP-equipped LLM produces precise, verifiable and appropriately quantified answers on 25/30 tasks vs 14/30 for web and 2/30 for bare (Wilcoxon p<0.01, Holm-corrected, all pairwise comparisons). The MCP advantage is largest for tasks where data quantification is required (89% vs 11% web). This work establishes MCP over ontology-backed knowledge graphs as an effective method to improve LLM response quality for neuroscience and connectomics data.

13
An Integrated Knowledge Graph and Network Medicine Pipeline for Drug Repurposing: Benchmarking Across Human Diseases and Application to Amyotrophic Lateral Sclerosis

Jiang, A.; Hu, J.; Abdulle, Y.; Pain, O.; Iacoangeli, A.

2026-07-08 bioinformatics 10.64898/2026.07.03.736387 medRxiv
Top 0.1%
1.0%
Show abstract

Drug repurposing offers a practical strategy to identify new therapeutic uses for approved drugs, potentially reducing the time and cost associated with conventional drug development. We present a novel three-stage drug repurposing pipeline that integrates knowledge graph-based gene prediction, network-based drug-disease association analysis, and systematic classification of candidate drugs by therapeutic class. The pipeline integrates DGLinker to predict novel disease-associated genes, SAveRUNNER to identify drug repurposing candidates, and ATC Category Enrichment Analysis (ATCEA) to prioritise candidates by pharmacological class. We benchmarked the pipeline across twelve diseases using DrugBank and MEDI2-HPS as validation resources. Utilising DGLinker-expanded disease-gene sets as input increased the number of predicted repurposed drugs, while overall discriminative performance remained stable across diseases (AUROC 0.71-0.77). Application of ATCEA consistently improved precision, F1-score, and specificity, while reducing recall, reflecting a conservative prioritisation strategy that contracts the candidate space while retaining pharmacologically coherent drug-disease candidates. We further applied the pipeline to amyotrophic lateral sclerosis (ALS), a neurodegenerative disease with limited therapeutic options, and performed a deeper literature-based validation of the results. Incorporation of DGLinker-predicted genes substantially increased the number of significant candidate drugs and uncovered enriched ATC categories not identified using known ALS genes alone, including antidepressants and antipsychotics. Moreover, several drugs with supporting evidence available in the literature were identified only when DGLinker-predicted genes were used. Overall, 77 candidate drugs were prioritised within significantly enriched ATC categories, several of which are supported by previously published studies. To provide exploratory real-world support for these findings, we further evaluated candidate drugs in a longitudinal electronic health record (EHR) dataset of 2361 patients with ALS from King's College Hospital. Although the number of evaluable drugs was limited due to sample size, the EHR analysis provided additional clinically relevant context for selected prioritised drugs and pharmacological classes. Our pipeline demonstrates potential to accelerate drug repurposing by integrating complementary computational approaches to each step of the process, providing an end-to-end framework that showed robust performance across benchmarking experiments and use cases.

14
Likelihood-Based Inference and Model Selection for Stochastic Gene Expression in Probability-Generating-Function Space

Wang, Y.; Shu, Z.; McAuley, K. B.; Cao, Z.

2026-08-25 systems biology 10.64898/2026.08.24.746673 medRxiv
Top 0.1%
1.0%
Show abstract

Selecting stochastic gene-expression models from single-cell counts requires accurate parameter inference and efficient model selection. Likelihood methods in count space can be costly when full stationary count distributions are unavailable, whereas approximate methods may lose accuracy. Probability generating functions (PGFs) offer a compact analytical alternative, but existing PGF workflows are generally not likelihood based and therefore rely on computationally intensive cross-validation. We develop a likelihood-based PGF framework for both tasks. Correlated empirical PGF values are used to construct a Gaussian quasi-likelihood for parameter inference and PGF-based Bayesian information criterion (BIC) for model selection. We show that the empirical PGF is exactly unbiased and that the parameter estimator is consistent, converges at the inverse-square-root sample-size rate, and is first-order asymptotically unbiased. For large samples and a uniquely preferred model, PGF-BIC selects the same model as leave-one-out cross-validation in PGF space.

15
HSSM: A Widely Applicable Toolbox for Hierarchical Bayesian Neuro-cognitive Modeling

Fengler, A.; Xu, Y.; Bera, K.; Paniagua, C.; Omar, A.; Frank, M. J.

2026-06-09 neuroscience 10.64898/2026.06.05.730398 medRxiv
Top 0.1%
0.9%
Show abstract

Computational models are central to cognitive neuroscience, but their rigorous application to experimental datasets is often constrained to a narrow set of canonical models that afford tractable analytical computations. We introduce the HSSM (Hierarchical Sequential Sampling Model) ecosystem, a Python toolbox that democratizes access to a broad, extensible array of neurocognitive process models through hierarchical Bayesian inference. Naturally leveraging simulation-based inference via likelihood surrogates, HSSM enables fast parameter estimation for models lacking closed-form likelihoods. Built atop PyMC and Bambi, HSSM provides a user-friendly formula syntax for specifying hierarchical mixed-effects regressions on model parameters, incorporating trial-by-trial neural or physiological covariates. The ecosystem allows fast model simulation and training data generation, as well as the neural network training utilities to deploy surrogate likelihood networks via HuggingFace. Contributions are designed to benefit not only the single researcher working on a problem, but organically, the entire research community. Together, the tools in the HSSM ecosystem bridge the interests of computational theorists as well as experimentalists, accelerating the cycle from model development to rigorous empirical testing.

16
Covariant Biochemical Systems Theory: cBST1~cBST3 Descriptors and Quantitative Validation

Oosawa, C.

2026-07-29 systems biology 10.64898/2026.07.26.732504 medRxiv
Top 0.1%
0.9%
Show abstract

Biochemical Systems Theory (BST) represents nonlinear biochemical rate laws by local power-law approximations in logarithmic concentration coordinates. First-order coefficients are elasticities, whereas higher-order derivatives describe local log-synergism and its variation. Ordinary higher derivatives, however, are not tensorial under nonlinear reparameterizations and can mix biochemical response structure with coordinate artifacts. We formulate a covariant hierarchy, cBST1-cBST3, on a positive operating-point space equipped with a declared reference connection. cBST1 recovers classical elasticities in a flat logarithmic chart, cBST2 is the covariant Hessian of the log-response, and cBST3 is the symmetrized covariant derivative of cBST2. The framework is quantitatively evaluated using three representative rate laws from the curated yeast glycolysis model BIOMD0000000064: glucose transport, glucose phosphorylation, and phosphofructokinase. For 10,000 finite log-concentration perturbations at each of five radii, cBST2 reduced the cBST1 log-rate root-mean-square error by 96.6-98.6% at the largest tested radius, and cBST3 provided a further 68.4-98.1% reduction. The contracted cBST2 and cBST3 terms strongly predicted the corresponding lower-order truncation errors. Under the nonlinear transformation qi = sinh(ui), covariant contractions agreed across coordinates to within a 95th-percentile relative error of 1.2 x 10-13, whereas ordinary higher derivatives showed order-unity coordinate mismatches. Supplementary analytic tests recovered the expected second-, third-, and fourth-order truncation-error scaling. These results show that cBST1-cBST3 are not only coordinate-consistent descriptors but also practical diagnostics of where local power-law approximations require higher-order correction.

17
Methodological guidelines for circadian modeling of Daylight Saving Time: application to the United States

Martin-Olalla, J. M.; Mira, J.

2026-06-22 public and global health 10.64898/2026.06.17.26355889 medRxiv
Top 0.1%
0.9%
Show abstract

Modeling the circadian impact of seasonal clock changing requires precise synchronization between solar and social time. This report critiques a recent study that associated disease prevalence in the United States with seasonal clock exposure. We identify a fundamental computational error in which a sign reversal of the longitudinal offset effectively inverted the US East-West axis, cross-correlating local health data with the circadian burden of hypothetical locations on the opposite side of a time zone. We outline the methodology for a correct modelization of the circadian process in the context of US geography.

18
Slow stress-load accumulation dominates BDNF-dependent gain in a ten-state computational model of stress biochemistry

Ahmad, I.

2026-07-21 systems biology 10.64898/2026.07.15.738784 medRxiv
Top 0.1%
0.9%
Show abstract

BackgroundAcute stress responses are often reversible, whereas sustained stress can produce coordinated disruption across endocrine, metabolic, inflammatory, antioxidant, and neuroplastic pathways. The Tiered Stress Biochemistry Model (TSBM) is a hypothesis-generating ten-state ordinary differential-equation framework linking a stylized cortisol signal to noradrenergic drive, vitamin C, a phenomenological aldosterone/renin-angiotensin drive, magnesium, normalized BDNF-related and Nrf2-related states, inflammation, and tryptophan-kynurenine metabolism. MethodsFive prespecified scenarios (normal, acute, chronic, depression-like, and low-cortisol PTSD-like) were simulated for 168 hours. Analyses included local stability, output-specific sensitivity screening, structural ablation, a 400-draw Latin-hypercube scan over independent {+/-} 20% parameter ranges, uncertainty distributions for threshold-crossing times, a success-conditioned parameter-trade-off screen, and a wider scan in which hypothesized BDNF-feedback gain magnitudes varied log-uniformly from 0.1 to 10 times nominal. Parameters were classified as literature-derived, literature-constrained/model-defined, calibrated, or hypothesized. ResultsUnder the specified forcing assumptions, sustained stress produced coordinated changes across several pathways. In the depression-like scenario, removing slow stress-load accumulation increased day-7 BDNF-related activity from 47.2% to 82.6%, reduced inflammation from 14.63 to 2.15 arbitrary units, and lowered KYN/TRP from 0.173 to 0.057. Removing BDNF-dependent gain increased BDNF only to 49.1% and delayed KYN/TRP crossing by 2.5 hours. The stricter multi-output conclusion was retained in 91% of uncertainty draws. Median crossing times retained the nominal sequence, but the complete four-event order occurred in only 43% of all draws. ConclusionsWithin this reduced model, a shared slow stress-load process coordinates the high-exposure state, whereas BDNF-dependent feedback acts mainly as an amplifier. The model generates an experimentally testable staging hypothesis: under sustained high-exposure forcing, magnesium changes may precede later BDNF-related and KYN/TRP changes. Longitudinal studies are required to determine whether this sequence occurs biologically, whether it is reversible, and whether it has clinical relevance. The simulations are not diagnostic or treatment recommendations.

19
Regulatory Bias Constrains Epigenetic Aging Trajectories

Morozova, T.; Polster, A.; Axelson-Fisk, M.

2026-07-23 systems biology 10.64898/2026.07.22.740004 medRxiv
Top 0.1%
0.9%
Show abstract

Aging reflects both stochastic fluctuation and biological regulation. We present a Markov chain framework for epigenetic aging that extends noise-driven models by adding a state-dependent bias term representing regulatory constraint. Using DNA methylation data from mice, rats, and bats, we show that empirical epigenetic aging is characterized by a progressive restriction of the accessible state space. We demonstrate that a stochastic model incorporating state-dependent regulatory bias successfully reproduces this constraint, whereas standard noise-driven models fail to capture it. Subsequently, the results indicate that a regression-based drift model can be used to predict future trajectories and achieve lower mean, covariance, and state-increment dependence errors than the biased model. The pattern of state loss is consistent with discrete bifurcation events, suggesting resilience declines stepwise rather than continuously. This implies that the solution space for intervention narrows irreversibly at each transition, making intervention timing critical.

20
Autonomous Spatial Transcriptomics Analysis (ASTA): Demonstrating Performance Improvements through Clustering, Biological Annotation, and AI-Driven Discovery

Zhang, M.; Roe, M.; Pollett, C.; Andreopoulos, W. B.

2026-08-18 bioinformatics 10.64898/2026.08.10.743848 medRxiv
Top 0.1%
0.8%
Show abstract

Spatial transcriptomics keeps measurement of gene expression while preserving spatial context, yet traditional analysis methods face challenges in computational efficiency, biological interpretability, and autonomous discovery. This project presents a framework solving these issues through three parts: (1) an ensemble clustering system achieving 66.7% improvement over baseline average and 23.9% over best single method with silhouette score of 0.540 and statistical significance (p = 0.0032, Cohens d = 1.82); (2) a knowledge-based clustering framework that annotates 88.6% of cells across 8 ovarian cell types using 428 marker genes; and (3) a GPT-4o-mini-powered autonomous agent that generated 3 biological hypotheses with validations.